Papers with statistical method

4 papers
Mandarinograd: A Chinese Collection of Winograd Schemas (2020.lrec-1)

Copied to clipboard

Challenge: Mandarinograd is a corpus of Winograd Schemas in Mandarin Chinese . WS are hard to collect and few datasets are publicly available .
Approach: They introduce a corpus of Winograd Schemas in Mandarin Chinese . they describe the difficulties faced when building the corpus and explain how they overcome the anomalies.
Outcome: The proposed corpus of Winograd Schemas in Mandarin Chinese is hard to build and resistant to statistical methods.
uniblock: Scoring and Filtering Corpus with Unicode Block Information (D19-1)

Copied to clipboard

Challenge: Existing methods to remove sentences consisting of illegal characters are tedious and repetitive.
Approach: They propose a statistical method to identify illegal characters in natural language processing . they use a fixed-size feature vector to generate a Gaussian mixture model for each sentence .
Outcome: The proposed method can score sentences and filter corpus on clean corpus and improve performance.
Stubborn Lexical Bias in Data and Models (2023.findings-acl)

Copied to clipboard

Challenge: Recent work has focused on spurious correlations between features and labels in training data . but, we find strong evidence of corresponding bias in the trained models .
Approach: They propose a method to reduce spurious correlations in training data by reweighting it using a large pool of extracted features.
Outcome: The proposed method reduces spurious correlations in training data, but still finds strong evidence of bias in trained models.
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show voice assistants do not perform equally well for everyone . however, research on demographic robustness of speech technologies is still scarce .
Approach: They propose a statistical method to detect demographic bias using a large dataset with controlled demographic tags.
Outcome: The proposed method shows statistically significant differences in performance across age, dialectal region and ethnicity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations